Add DSV4 GB200 Dynamo+SGLang AgentX recipes without W4A4 MegaMoE - #3629
Conversation
Remove enable-w4a4-mxfp4-megamoe from every prefill and decode worker in the six GB200 Dynamo+SGLang AgentX disaggregated recipes. The MegaMoE all-to-all backend and FP4 indexer settings are unchanged. Co-Authored-By: Claude Opus 5.5 <[email protected]>
Co-Authored-By: Claude Opus 5.5 <[email protected]>
|
Thanks for the contribution!
中文感谢你的贡献!
|
There was a problem hiding this comment.
Nothing blocking. The comments below are optional suggestions. There is no need to push a fix for them before merging.
Beyond the inline finding, I also checked the PR's core claim — that enable-w4a4-mxfp4-megamoe was removed from all prefill/decode workers in the six disaggregated recipes while moe-a2a-backend: megamoe stays — and confirmed no enable-w4a4-mxfp4-megamoe key remains anywhere in the new recipe files, and container.image matches the master-config image field consistently across all eight new scenarios.
Extended reasoning...
The diff adds 7 new srt-slurm recipe YAMLs plus matching master-config entries, a runners.yaml model-path mapping, a lanes.py policy-table entry, and an append-only changelog addition for a new DeepSeek-V4-Pro-0813 GB200 FP4 AgentX workload; no auth/crypto/permission code is touched, only benchmark config and launch policy tables. A confirmed finding already flags a real routing problem in runners.yaml (new model lacks a numa1-local alias, so it is not visible for node-local staging as the changelog implies), which is sufficient reason a human should look before merge.
|
View unofficial run (performance): https://inferencex.semianalysis.com/inference?unofficialRun=36833446190 View unofficial run (accuracy): https://inferencex.semianalysis.com/evaluation?unofficialRun=36833446190 |
|
/use 36833446190 Reuse the passing full sweep for this PR: attempt 1, tested commit 中文复用本 PR 已通过的完整 sweep:attempt 1,测试 commit |
|
As a PR reviewer and CODEOWNER, I have reviewed this and have:
Additional detail section:
Signed: |
Ankur-singh
left a comment
There was a problem hiding this comment.
Reviewed the benchmark and eval evidence for 2a58623b5a441f1b5c190ba9666a70881301d1d3. CODEOWNER checklist. The engine-first item remains unchecked pending core-maintainer resolution. Merge conflicts also remain.
中文
已审阅 2a58623b5a441f1b5c190ba9666a70881301d1d3 的 benchmark 和 eval 证据。CODEOWNER checklist。engine-first 项仍未勾选,等待核心维护者确认;merge conflicts 也仍待解决。
✅✅✅ Verdict: PASS ✅✅✅Assessed pinned head Passed and not applicable checks✅ Check 0 (CODEOWNER): PASS — ✅ Check 1 (Passing sweep on in-PR commit): PASS — On pinned head ✅ Check 2 (Evals pass): PASS — ➖ Check 3 (Upstream recipe): N/A — disaggregated/multi-node submission; the recipe-link requirement applies to single-node recipes only. All 8 recipes are under ✅ Check 4 (Reuse command): PASS — ✅ Check 5 (Latest checklist template): PASS — All 17 items in the current ✅ Check 6 (Upstream images / engine-first): PASS — Both new entries use upstream ✅ Check 7 (No deprecated models/scenarios): PASS — ✅ Check 8 (No architecture hacks / event publication): PASS — There are no ✅ Check 9 (Spec-decode via chat template): PASS — The logged aiperf command uses ✅ Check 10 (No engine patches): PASS — The diff has no ✅ Check 11 (Agentic golden AL): PASS — ➖ Check 12 (Append-only): N/A — Neither new ✅ Check 13 (Draft runs as shipped): PASS — The draft is the bundled DSpark head of ✅ Check 14 (Pareto coverage): PASS — One curve: dsv4 / agentic-coding / GB200 dynamo-sglang / fp4 / P90 E2EL / run 36833446190 / Assessed commit: |
| shared_run_root=( | ||
| Match(any_of("minimaxm3", "kimik3", "qwen3.5", "glm5.2")), | ||
| Match(any_of("dsv4"), frameworks=any_of("dynamo-vllm")), | ||
| Match(any_of("dsv4"), frameworks=any_of("dynamo-sglang"), agentic=True), |
There was a problem hiding this comment.
@adibarra can you take a look since change to infx
Preserve main changelog entries and append the GB200 DSV4 contribution.
Add DSV4-Pro GB200 Dynamo+SGLang AgentX recipes using DSpark block size 6 and HiCache on
lmsysorg/sglang:nightly-dev-20260916-c9a8fba9.Coverage: TP8 aggregate at C1/C4; 1P1D DEP8/DEP16 at C64/C128; 1P1D DEP16/DEP32 at C256; and 2P1D DEP16/DEP32 at C768/C1024/C1280.
The disaggregated recipes omit
enable-w4a4-mxfp4-megamoe, retain the MegaMoE all-to-all backend and FP4 indexer, and use decode request caps of 512/1024 for DEP16/DEP32. C128 prefill usesmem-fraction-static: 0.80.